Papers with low-resource setting

43 papers
Handling Noisy Labels for Robustly Learning from Self-Training Data for Low-Resource Sequence Labeling (N19-3)

Copied to clipboard

Challenge: In low-resource environments, self-training is less effective due to unreliable annotations . we combine self-teaching with noise handling to clean the self-labeled data .
Approach: They propose to combine self-training with noise handling to clean unlabeled data . they propose to model clean and noisy labels separately to improve performance .
Outcome: The proposed method performs better than baseline methods on Chunking and NER.
Mention Flags (MF): Constraining Transformer-based Text Generators (2021.acl-long)

Copied to clipboard

Challenge: Constrained decoding algorithms produce hypotheses satisfying all constraints, but they are computationally expensive and can lower the generated text quality.
Approach: They propose a Mention Flag mechanism which traces whether lexical constraints are satisfied in outputs of an S2S decoder.
Outcome: The proposed models maintain higher constraint satisfaction and text quality than baseline models and other constrained decoding algorithms.
Making a Point: Pointer-Generator Transformers for Disjoint Vocabularies (2020.aacl-srw)

Copied to clipboard

Challenge: Existing neural models rely on an overlap between source and target vocabularies to perform sequence-to-sequence tasks.
Approach: They propose a pointer-generator transformer model for disjoint vocabularies that does not rely on an overlap between source and target vocs.
Outcome: The proposed model outperforms a standard pointer-generator transformer by an average of 5.1 WER over 15 languages.
A Little Pretraining Goes a Long Way: A Case Study on Dependency Parsing Task for Low-resource Morphologically Rich Languages (2021.eacl-srw)

Copied to clipboard

Challenge: Neural dependency parsing has been a success for many domains and languages, but the bottleneck of massive labelled data limits its effectiveness for low resource languages.
Approach: They propose to use morphological knowledge to improve dependency parsing for morphology rich languages in a low-resource setting to perform experiments.
Outcome: The proposed method achieves an average gain of 2 points (UAS) and 3.6 points (LAS) on 10 MRLs in low-resource settings.
Neural Unsupervised Parsing Beyond English (D19-61)

Copied to clipboard

Challenge: Unsupervised parsing is a task that can be learned without substantial prior knowledge.
Approach: They train an unsupervised model for Arabic, Chinese, English, and German to learn syntactic structure from unlabeled text.
Outcome: The PRPN architecture outperforms trivial baselines and acquires at least some parsing ability for all languages.
Sequence Tagging with Contextual and Non-Contextual Subword Representations: A Multilingual Evaluation (P19-1)

Copied to clipboard

Challenge: Pretrained contextual and non-contextual subword embeddings are available in over 250 languages, allowing massively multilingual NLP.
Approach: They compare pretrained contextual and non-contextual subword embeddings with a contextual representation method, namely BERT, on multilingual named entity recognition and part-of-speech tagging.
Outcome: The proposed method outperforms non-contextual embeddings on multilingual named entity recognition and part-of-speech tagging.
Towards Efficient Dialogue Processing in the Emergency Response Domain (2023.acl-srw)

Copied to clipboard

Challenge: Adapters perform dialogue act classification and domain-specific slot tagging in the emergency response domain.
Approach: They propose to build a system that performs dialogue act classification and domain-specific slot tagging while being efficient, flexible and robust.
Outcome: The proposed model performs well in the emergency response domain while being efficient, flexible and robust.
Compositional Representation of Morphologically-Rich Input for Neural Machine Translation (P18-2)

Copied to clipboard

Challenge: Neural machine translation models are typically trained with fixed-size input and output vocabularies, which creates a bottleneck on their accuracy and generalization capability.
Approach: They propose to replace the source-language embedding layer of NMT with a bi-directional recurrent neural network that generates compositional representations of the input at any desired level of granularity.
Outcome: The proposed approach outperforms existing methods in a low-resource setting with five languages . the proposed approach consistently outperformed existing methods with a single word representation .
Improving Cross-lingual Transfer through Subtree-aware Word Reordering (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent studies show that multilingual language models are not effective when dealing with less-represented languages.
Approach: They propose a powerful reordering method that learns word-order patterns conditioned on the syntactic context from a small amount of annotated data.
Outcome: The proposed method outperforms baselines on a variety of tasks and is effective in both zero-shot and few-shot scenarios.
Building an Efficient Multilingual Non-Profit IR System for the Islamic Domain Leveraging Multiprocessing Design in Rust (2024.emnlp-industry)

Copied to clipboard

Challenge: Existing models that are pre-trained on a general domain can deteriorate performance due to domain shift when applied to new domains.
Approach: They propose to train a multilingual non-profit IR system for the Islamic domain using Rust Language capabilities.
Outcome: The proposed model outperforms models pre-trained on general domains and on resource-constrained devices.
One More Modality: Does Abstract Meaning Representation Benefit Visual Question Answering? (2025.findings-emnlp)

Copied to clipboard

Challenge: incorporating explicit semantic information, in the form of Abstract Meaning Representation graphs, can enhance VQA models.
Approach: They augment two vision-language models with sentence- and document-level AMRs . they find that in well-resourced settings, models are negatively impacted by AMR .
Outcome: The proposed model improves in well-resourced and low-resource settings with AMR graphs . the model achieves 13.1% relative gain using sentence-level AMRs compared with the smaller model .
Generative Data Augmentation for Commonsense Reasoning (2020.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in commonsense reasoning depend on large-scale human-authored training data.
Approach: They propose a generative data augmentation technique that augments human-authored training data by using pretrained language models.
Outcome: The proposed technique outperforms existing methods on commonsense reasoning benchmarks and enhances out-of-distribution generalization.
Global Structure Knowledge-Guided Relation Extraction Method for Visually-Rich Document (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods focus on manipulating entity features to find pairwise relations, yet neglect the more fundamental structural information that links disparate entity pairs together.
Approach: They propose a Visual Relation Extraction framework that generates relation predictions on entity pairs extracted from scanned images and incorporates global structural knowledge into the representations of the entities.
Outcome: The proposed framework outperforms existing methods in fine-tuning setting and yields stronger data-efficient performance in the low-resource setting.
Extractive Summarization of Legal Decisions using Multi-task Learning and Maximal Marginal Relevance (2022.findings-emnlp)

Copied to clipboard

Challenge: Summarizing legal decisions requires the expertise of law practitioners, which is time- and cost-intensive.
Approach: They propose methods for extracting summarized legal decisions using limited expert annotated data.
Outcome: The proposed models achieve ROUGE scores vis-à-vis expert extracted summaries that match inter-annotator comparisons.
Better Character Language Modeling through Morphology (P19-1)

Copied to clipboard

Challenge: Inflected words benefit more from explicitly modeling morphology than uninflectes . morphological supervision is also used to augment character language models in low-resource languages .
Approach: They add morphological supervision to character language models via multitasking to improve BPC performance across 24 languages even when morphology data and language modeling data are disjointed.
Outcome: The addition improves performance even when morphology data and language modeling data are disjointed.
A Three-Stage Learning Framework for Low-Resource Knowledge-Grounded Dialogue Generation (2021.emnlp-main)

Copied to clipboard

Challenge: Existing knowledge-grounded dialogues perform poorly when transfer into new domains with limited training samples.
Approach: They propose a weakly supervised three-stage learning framework based on weakly-supervised learning based upon large scale ungrounded dialogues and unstructured knowledge base.
Outcome: The proposed framework outperforms state-of-the-art methods even in zero-resource setting.
Phone Features Improve Speech Translation (2020.acl-main)

Copied to clipboard

Challenge: End-to-end models for speech translation more tightly couple speech recognition (ASR) and machine translation (MT) compared to cascades, but performance gap remains in low-resource conditions .
Approach: They propose two methods to incorporate phone features into current neural speech translation models.
Outcome: The proposed models outperform existing models and cascades by up to 9 BLEU on low-resource conditions.
Towards Summarizing Healthcare Questions in Low-Resource Setting (2022.coling-1)

Copied to clipboard

Challenge: Existing methods to generate large-scale datasets are difficult in closed domains where human annotation requires domain expertise.
Approach: They propose a method to generate diverse and semantic questions in a low-resource setting with the aim of summarizing healthcare questions.
Outcome: The proposed method generates diverse, fluent, and informative summarized questions on healthcare question summarization datasets.
Normalizing Compositional Structures Across Graphbanks (2020.coling-main)

Copied to clipboard

Challenge: Graph-based meaning representations (MRs) exhibit structural differences that reflect different theoretical and design considerations, presenting challenges to uniform linguistic analysis and cross-framework semantic parsing.
Approach: They propose a method to normalize MRs at the compositional level by linguistically-grounded rules.
Outcome: The proposed method increases the match in compositional structure between MRs and improves multi-task learning in a low-resource setting.
Translation via Annotation: A Computational Study of Translating Classical Chinese into Japanese (2026.eacl-long)

Copied to clipboard

Challenge: Ancient people translated classical Chinese into Japanese using a system of annotations placed around characters.
Approach: They propose to introduce an LLM-based annotation pipeline and construct a dataset from digitized open-source translation data to improve sequence tagging tasks.
Outcome: The proposed method achieves high scores on direct machine translation, but could serve as a supplement to LLMs to improve the quality of character’s annotation.
Use Random Selection for Now: Investigation of Few-Shot Selection Strategies in LLM-based Text Augmentation (2025.findings-emnlp)

Copied to clipboard

Challenge: generative large language models are increasingly used for data augmentation tasks . text samples are mostly selected randomly and a comprehensive overview of other sample selection strategies is lacking.
Approach: They compare random sample selection strategies and random sample sampling strategies to evaluate their effects in a low-resource setting.
Outcome: The proposed model performance improvements are compared with other sample selection strategies.
Cross-language Sentence Selection via Data Augmentation and Rationale Training (2021.acl-long)

Copied to clipboard

Challenge: a new approach to cross-language sentence selection is proposed for low-resource contexts . a cross-lingual embedding-based model is proposed that avoids translation entirely .
Approach: They propose a cross-lingual embedding-based query relevance model that uses data augmentation and negative sampling techniques to directly learn a query-sentence pair.
Outcome: The proposed approach performs better than state-of-the-art models on noisy parallel data . consistent improvements are seen across three language pairs over state- of-the art models .
Self-supervised Graph Masking Pre-training for Graph-to-Text Generation (2022.emnlp-main)

Copied to clipboard

Challenge: Large-scale pre-trained language models (PLMs) have advanced Graph-to-Text generation by processing the linearised version of a graph.
Approach: They propose to mask pre-training tasks that neither require supervision signals nor adjust the architecture of the underlying pre-trained encoder-decoder model.
Outcome: The proposed method achieves state-of-the-art results on WebNLG+2020 and EventNarrative datasets and is very efficient in the low-resource setting.
CodePrompt: Task-Agnostic Prefix Tuning for Program and Language Generation (2023.findings-acl)

Copied to clipboard

Challenge: Prompt-tuning methods have been used to solve inefficient parameter update and storage issues in Natural Language Generation tasks.
Approach: They propose a task-agnostic prompt tuning method that reflects the traits of PLM for program language.
Outcome: The proposed method is effective in three PLG tasks, not only in the full-data setting but also in the low-resource setting and cross-domain setting.
Automatic Readability Assessment for Closely Related Languages (2023.findings-acl)

Copied to clipboard

Challenge: In recent years, the main focus of research on automatic readability assessment (ARA) has shifted towards using expensive deep learning-based methods with the primary goal of increasing models’ accuracy.
Approach: They focus on how linguistic aspects such as mutual intelligibility or degree of language relatedness can improve ARA in a low-resource setting.
Outcome: The inclusion of CrossNGO, a novel feature exploiting n-gram overlap, significantly improves the performance of ARA models compared to the use of off-the-shelf large multilingual language models alone.
Mask-then-Fill: A Flexible and Effective Data Augmentation Framework for Event Extraction (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing data augmentation methods for event extraction are costly and time-consuming.
Approach: They propose a data augmentation framework that randomly masks out an adjunct sentence fragment and infills a variable-length text span with a fine-tuned infilling model.
Outcome: The proposed framework can generate more diverse data while keeping the original structure unchanged . it can replace a fragment of arbitrary length in the text with another fragment of variable length .
Multilingual LLMs are Better Cross-lingual In-context Learners with Alignment (2023.acl-long)

Copied to clipboard

Challenge: a handful of studies have explored ICL in a cross-lingual setting . emergence of large-scale, pretrained, Transformer-based language models has marked the commencement of an avant-garde era in NLP.
Approach: They propose a novel prompt construction strategy to bridge the gap between ICL and cross-lingual text classification.
Outcome: The proposed approach outperforms random prompt selection by a large margin across three tasks using 44 different cross-lingual pairs.
Meta-Transfer Learning for Code-Switched Speech Recognition (2020.acl-main)

Copied to clipboard

Challenge: Increasing number of people in the world today speak a mixed-language as a result of being multilingual.
Approach: They propose a method to transfer learn on a code-switched speech recognition system by extracting information from high-resource monolingual datasets.
Outcome: The proposed model outperforms baselines on speech recognition and language modeling tasks and is faster to converge.
Effectiveness of Data Augmentation and Pretraining for Improving Neural Headline Generation in Low-Resource Settings (2022.lrec-1)

Copied to clipboard

Challenge: Neural approaches for natural language generation (NLG) have mushroomed due to large textual resources.
Approach: They propose to use a pretrained multilingual encoder-decoder model and a combination of two pretrained language models to train a model in a low-resource setting.
Outcome: The proposed model outperforms the previous model on English and on a small subset of the same data.
Constrained Labeled Data Generation for Low-Resource Named Entity Recognition (2021.findings-acl)

Copied to clipboard

Challenge: Named Entity Recognition (NER) in lowresource languages has been a challenge for years . Existing methods suffer from low quality of annotated data in target language .
Approach: They propose a method that uses projected annotations to generate pseudo supervised data with a transformer language model and a constrained beam search.
Outcome: The proposed method achieves state-of-the-art or competitive performance in low-resource languages.
Data Augmentation for Context-Sensitive Neural Lemmatization Using Inflection Tables and Raw Text (N19-1)

Copied to clipboard

Challenge: Using context-sensitive approaches to lemmatization can improve accuracy on unseen and unseense words.
Approach: They propose to use inflection tables and Wikipedia sentences to train a lemmatizer with little or no labeled corpus data to combine type-based learning with context.
Outcome: The proposed model generalizes from unambiguous examples, improving overall and especially on unseen words.
Tackling the Low-resource Challenge for Canonical Segmentation (2020.emnlp-main)

Copied to clipboard

Challenge: morphological segmentation is a task of dividing words into their constituting morphemes . we compare two new approaches for the task when training data is limited .
Approach: They propose to use an LSTM pointer-generator and a sequence-to-sequence model to perform canonical segmentation when training data is limited.
Outcome: The proposed models outperform existing models on German, English, and Indonesian in low-resource scenarios by 11.4% accuracy.
ParaMac: A General Unsupervised Paraphrase Generation Framework Leveraging Semantic Constraints and Diversifying Mechanisms (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing unsupervised methods for paraphrase generation are weak in semantic equivalence or expression diversity.
Approach: They propose a framework for unsupervised paraphrase generation that employs multi-aspect equivalence constraints and multi-granularity diversifying mechanisms to achieve good semantic equvalence and expressive diversity.
Outcome: The proposed framework achieves 9.1% and 3.3% absolute gains over previous SOTA on Quora and MSCOCO and can improve to 18.0% and 4.6% on GLUE.
AdaptSum: Towards Low-Resource Domain Adaptation for Abstractive Summarization (2021.naacl-main)

Copied to clipboard

Challenge: State-of-the-art abstractive summarization models rely on extensive labeled data, which lowers their generalization ability on domains where such data are not available.
Approach: They propose to use domain adaptation methods to simulate the low-resource domain adaptation setting for abstractive summarization systems with existing datasets across six diverse target domains.
Outcome: The proposed model can be used to adapt to a low-resource domain adaptation setting.
Separating Context and Pattern: Learning Disentangled Sentence Representations for Low-Resource Extractive Summarization (2023.findings-acl)

Copied to clipboard

Challenge: Context information is one of the key factors for extractive summarization, but other factors can be used to identify sentence importance.
Approach: They propose to disentangle context and pattern factors for extractive summarization . they separate context and patterns for a better generalization ability in low-resource setting .
Outcome: The proposed model can be used in the zero-shot setting or fine-tuned in the few-shot settings.
LVLM-Aware Multimodal Retrieval for RAG-Based Medical Diagnosis with General-Purpose Models (2026.findings-acl)

Copied to clipboard

Challenge: Using retrieval augmentation, large vision language models can be used for diagnostic accuracy, but multimodal retrieval-augmented diagnosis is challenging.
Approach: They propose a lightweight mechanism for enhancing diagnostic performance of retrieval-augmented LVLMs by fine-tuning a multimodal retriever and general-purpose backbone models.
Outcome: The proposed mechanism achieves competitive results without medical training compared to pre-trained models with extensive training.
Cross-Lingual Abstractive Summarization with Limited Parallel Resources (2021.acl-long)

Copied to clipboard

Challenge: Existing approaches to cross-lingual summarization use limited available cross-linguistic resources.
Approach: They propose a multi-task framework for cross-lingual abstractive summarization that uses a single decoder to generate monolingual and cross-linguistic summaries.
Outcome: Experiments on two CLS datasets show that the proposed model outperforms baseline models in low-resource and full-dataset scenarios.
CASSI: Contextual and Semantic Structure-based Interpolation Augmentation for Low-Resource NER (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for text augmentation suffer from annotation corruption for token-level tasks like NER.
Approach: They propose a novel augmentation scheme that generates high-quality contextually diverse augmentations while avoiding annotation corruption.
Outcome: The proposed scheme outperforms existing methods at multiple low resource levels, in multiple languages, and for noisy and clean text.
Graph-Induced Transformers for Efficient Multi-Hop Question Answering (2022.emnlp-main)

Copied to clipboard

Challenge: Recent MHQA tasks that require inter-paragraph/sentence linkages use graphs to model internal structural information within text.
Approach: They propose a graph-induced transformer that applies graph-derived attention patterns directly into a PLM without external graph modules.
Outcome: The proposed model can replace external graph modules while preserving model performance.
BanglaAbuseMeme: A Dataset for Bengali Abusive Meme Classification (2023.emnlp-main)

Copied to clipboard

Challenge: a number of studies have tried to detect and control the spread of such abusive memes on social media platforms.
Approach: They build a Bengali meme dataset to test models for abusive memes . they find that multimodal models that use both textual and visual information outperform unimodal models .
Outcome: The proposed model outperforms unimodal models in a Bengali meme dataset.
Towards Low-Resource Alignment to Diverse Perspectives with Sparse Feedback (2025.findings-emnlp)

Copied to clipboard

Challenge: popular training paradigms for language models often assume there is one optimal answer for every query.
Approach: They propose to enhance pluralistic alignment of language models using pluralistic decoding and model steering methods.
Outcome: The proposed methods improve pluralistic alignment of language models in a low-resource setting . the proposed methods decrease false positives in several high-stakes tasks .
Statistical and Neural Methods for Hawaiian Orthography Modernization (2025.emnlp-main)

Copied to clipboard

Challenge: Hawaiian orthography employs two distinct spelling systems, both of which are used by communities of speakers today.
Approach: They develop models that convert between the ‘okina letter and kahak diacritic, which represent glottal stops and long vowels, respectively.
Outcome: The proposed models outperform neural seq2seq models and LLMs in a low-resource setting, highlighting the potential for traditional machine learning approaches in . low-cost environments.
PO-KGQA: Preference Optimization for Low-Resource Complex Knowledge Graph Question Answering (2026.findings-acl)

Copied to clipboard

Challenge: Existing low-resource in-context learning-based knowledge graph question answering methods rely heavily on large language models to convert natural language questions into logical forms.
Approach: They propose a low-resource in-context learning-based knowledge graph question answering (KGQA) that uses large language models to convert a natural language question into its corresponding logical form.
Outcome: The proposed method outperforms other methods on complex benchmarks by approximately 9% (avg).

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations